Avatar of Shangeth Rajaa

Shangeth Rajaa

Open to full-time & consulting roles in Voice AI

Senior Research Scientist working on Voice AI, Turn-Taking, Full-Duplex Spoken Dialogue Systems, and Multi-Modal Speech LLMs.

Resume
  • About
  • CV
  • Publications
  • Blog
  • Courses

#speech llm

Content tagged with "speech llm"

DualTurn: Learning Turn-Taking from Dual-Channel Generative Speech Pretraining
2026-04-03
#Voice AI #Turn-Taking #Full-Duplex #Spoken Dialogue #Speech LLM

Speech-to-speech models know when to speak but can't reason. Cascaded LLM pipelines can reason but only react to silence. DualTurn pretrains on dual-channel human conversation to bring S2S-level turn-taking into a standard ASR-LLM-TTS stack.

DualTurn: Learning Turn-Taking from Dual-Channel Generative Speech Pretraining
2026-03-09 Shangeth Rajaa Interspeech 2026 (Accepted)
#Voice AI #Turn-Taking #Spoken Dialogue #Speech LLM

Dual-channel generative pretraining for learning natural turn-taking in spoken dialogue without labeled data. A 0.5B model that outperforms models 6x its size on turn prediction.

View
SpeechLLM: Multi-Modal LLM for Speech Understanding
2024-06-26
#Speech LLM #Voice AI #Speech Representation #Prosody

A small multimodal LLM that reads paralinguistic signal — emotion, prosody, speaker traits — directly from speech audio instead of through an ASR transcript, built alongside the release of SpeechLLM at Skit.ai.

Speech LLMs for Conversations
2024-05-09
#Voice AI #Speech LLM #Conversational AI

A multimodal speech LLM that processes audio directly to enhance conversational AI while reducing overhead compared to traditional ASR-LLM-TTS pipelines.

View
© 2026 Shangeth Rajaa.